MacOS 26.7 Tahoe Release Candidate Contains a Video Demonstrating Camera-Equipped AirPods in Action 

Oops.


Follow-Up Thoughts on Watermarking Schemes for AI-Generated Text

Some follow-up to this weekend’s stemwinder “Anthropic’s ‘Watermark’ Text Adulteration in Claude Is a Perversion of Writing”:

Temperature

Contra a bunch of idiots at Hacker News and elsewhere, I understand that popular LLMs do not just pick the “best” token (word) at each decision point. Counterintuitively, always selecting the highest-probability option produces undesirable results. So the models apply some randomization, and “temperature” is the term for the weighting that’s applied so that the “better” (higher-ranked by the model) choices have a higher chance of being chosen.

With a temperature of 1, models use their built-in probability distribution. With a temperature greater than 1, this distribution gets flatter — less-likely alternatives get a higher probability of being selected, and more-likely alternatives lower. With a temperature lower than 1, the probability distribution leans more toward the higher-ranked options. And with a temperature of 0, the highest-ranked option is always chosen. A temperature of 0 generally produces undesirable results — too predictable, too likely to get stuck. Like over-smoothing an image from a camera sensor, eliminating all noise makes the overall result worse, even if each single bit of “noise”, evaluated in isolation, is in some sense wrong.

The temperature-based randomness — which is what makes LLM output non-deterministic — is in place to help make the output better. The prose is clearly better with a temperature of 1 (with weighted randomness) than at temperature 0 (with no randomness). The watermarking schemes, on the other hand, are applying predictable-with-the-secret-key randomness for an entirely different purpose than improving the quality of the output, and thus, I believe, inherently make the output at least slightly worse.

Advocates of LLM watermarking schemes for text argue that the schemes don’t necessarily lower the quality of the generated prose, because they don’t change the temperatures — they only change the source of the randomness. Daniel Jalkut wrote a good piece today about this. I hope that’s true. I believe it’s possible that it is true. I think it’s highly unlikely that it is true. I do not see how a detectable signal can be added encoded in the choice of words without affecting the meaning of the prose. If it were true I think they’d show examples proving that it’s true. Also, Anthropic itself admits that it can’t properly watermark text that is programming language code:

For the same reason, code — which in very many cases has to be exact — has generally less watermarking than some other forms of text.

Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code. But by definition, it will have a negligible effect on the actual code produced.

I hold that good prose is much more like programming code. Exactness in word choice, phrasing, tone, and even punctuation is always better than imprecision. The difference is that sloppy programming code doesn’t run, or doesn’t run correctly. The human brain, on the other hand, is adept at parsing and making sense out of inexact, even sloppy, prose.

I Object Even If Quality Isn’t Adversely Affected

I do not believe these schemes can work without degrading prose quality, if only slightly. Again, though, I am open to being proven wrong. But even if we concede for the moment that such watermarking schemes do not necessarily degrade the quality of generated prose — not one iota — I still object to their use when they are being applied secretly, behind users’ backs. A useful watermark would be one that anyone can check. These SynthID “watermarks” are entirely dependent upon secrets held by the LLM providers (so far, Anthropic/Claude and Google/Gemini). I find that unacceptable, for reasons I hopefully made clear in my essay.

The people in favor of this watermarking for text have been sold a pipe dream, a fantasy. I’ve encountered dozens of comments from angry AI haters (many of them on Bluesky in particular, but also Threads and Hacker News) who are convinced that the only people who could be against the watermarking of AI-generated text are those who are duplicitously passing off AI-generated text as their own writing — and thus that I must be upset only because the jig will soon be up for me too. This of course is not true. I don’t even use AI to write text messages or emails for me, let alone a single sentence of my work.

But I find it funny that so many people who claim to believe that LLMs only produce “slop” and never anything useful also seem 100 percent convinced that the same LLMs are capable of watermarking their output in reliable ways. These people so desperately want to be able to point a finger at AI-generated text that they’ve fallen hook, line, and sinker for the argument from Google and Anthropic that, thanks to them, they’ll be able to.

I don’t want to spend too much time thinking about this because it’s a waste of time, but how exactly do these people think the existence of these mandatory watermarks and detection tools will change anything for the better? Let’s say you work at an office and you suspect that numerous of your colleagues are using AI to write emails and other work-related messages. Their messages are too long, too prolific, and lack lucidity. What are you going to do now? Copy and paste each of their messages into the watermark detectors from Anthropic, Google, and OpenAI? There cannot exist a single detector for all LLMs. And even if you find out that it says it’s a match, that an email or blog post or Slack message was very likely generated by, say, Claude, what are you going to do? March into your colleague’s office and tell them you caught them?

Anyone in a situation where “getting caught” would matter — students, say — is going to use non-watermarking LLMs or run their watermarked text through paraphrasing tools like Declaude.

No practical good is going to come of this, even if these watermarking schemes work as promised (and to be clear, I don’t believe any of it is going to work as promised).1

My advice is not to care whether anything was written by an AI or a human. The only thing worth evaluating is what we human readers are naturally good at determining: whether it is good or bad. If it’s good, read it. If it’s not, don’t. If you’ve got a job where you’re surrounded by colleagues filling your inbox with AI-generated messages that you can’t abide, get a new job or learn to live with it. Hidden secret watermarking signals — even if they work — aren’t going to make things go back to the way they used to be. If you read something and enjoy it, and subsequently find out it was generated by an LLM, don’t feel bad. You read something good that you enjoyed.

I read something earlier today that claimed most of the posts on LinkedIn are generated by AI. That the whole platform is just inundated with AI slop. Maybe it is, but I wouldn’t know, because I never look at LinkedIn because it’s always been filled with crap. If it smells like crap it’s crap, whether the turds came out of a human anus or a turd-generating robot.

The Argument That Only People Can Truly Write

Dan Moren, writing at Six Colors today, “LLMs Aren’t Writing”:

LLMs do not care about the words that they pick because they cannot care about anything.

Speaking of two things that are not the same, John rightly points out the difference between the phrases “he leaped at the chance” and “he jumped at the opportunity”. Those are indeed distinct — if semantically similar — phrases, each of which might be more apt in a particular situation; or, to put it in another fashion: the use of each of those phrases tells us something different, whether about the person being described or the writer.

But the LLM doesn’t know which of those phrases is the right phrase to use. It has a guess, based on its models and weights and inputs. But the ultimate choice of those phrases tells us nothing about the writer because there is no writer.

Moren’s is a fine retort to my post, but I fundamentally disagree — albeit at a philosophical level. If you’re reading a written work only to gain insight into the mind that produced it, there is no mind on the other end of AI-generated text. But the work itself exists. My disagreement with Moren starts and effectively ends with his (wonderfully summative) headline. I say if you can read something, it was necessarily written.

Again, this is philosophical. Was a photorealistic image generated by AI photographed? No, I would say it was not. Photography, I would say, is the act of focusing light through a lens onto a capturing sensor, capturing, to some extent, reality. I think Moren is arguing that writing is like that. If photography captures a physical scene from reality, writing captures thoughts from an actual mind. That something you can read that was produced by an LLM was merely generated in a way that doesn’t qualify as writing. Semantics. I just care about the article of text. Moren argues that LLMs are not writing; I say they are. But we’re disagreeing only over what the word writing means, not what is being produced.

As for “caring” about the difference between semantically similar but tonally different phrases, like “he leaped at the chance” versus “he jumped at the opportunity”, no, of course the LLM doesn’t “care”. But I, the reader, care very much. I wrote a column back in November on ChatGPT changing (and renaming) the “personalities” it allows users to choose from. These personalities generate text with strikingly different styles and tones. Because I use ChatGPT, I care very much about the tone and style of its responses to my queries. Not because I’m ever going to pass them off as my own writing, but because I’m the one who is reading them.

Moren, near the end of his column:

In the end, I can’t summarize it any better than to ask: if you care so much about word choice, why are you using AI to generate text?

If this does truly make AI-generated text worse, well… good. A lot of people are already willing to accept what an LLM churns out as “good enough” and, if I’m being realistic, I don’t think this will change anything. But if it does lead to more people being dissatisfied with the pablum they’re being fed and turning instead to writing and editing their own text, then that would actually be a positive outcome. Maybe it’d even mean fewer human writers being put out of jobs.

I sympathize, but I must disagree that it can possibly be seen as a net good for LLMs to produce worse prose. I read the output of LLMs every day. I use AI to generate text because I ask it questions (in text). I want the answers that I read to be cogent, lucid, accurate, blessedly terse — and ideally to strike a consistent tone that is pleasant to my reading ear. The genie is not going back in the bottle.

English Is the Finest Language, and Thus, Perhaps, More Fingerprintable

Lastly, here’s an interesting point to ponder. English is the most expressive language in the world. Don’t take my word for it — it’s the only language I speak (despite four years of Spanish in high school). Take the word of famed 20th century author Jorge Luis Borges, an Argentine polyglot whose first language was Spanish. In 1977 he was the guest on William F. Buckley’s “Firing Line”. You can (and should) watch the interview on YouTube, but here’s a transcript of the relevant portion from Jordan M. Poss:

Borges: I have done most of my reading in English. I find English a far finer language than Spanish.

Buckley: Why?

Borges: Well, many reasons. Firstly, English is both a Germanic and a Latin language. Those two registers — for any idea you take, you have two words. Those words will not mean exactly the same. For example if I say “regal” that is not exactly the same thing as saying “kingly.” Or if I say “fraternal” that is not the same as saying “brotherly.” Or “dark” and “obscure.” Those words are different. It would make all the difference — speaking for example — the Holy Spirit, it would make all the difference in the world in a poem if I wrote about the Holy Spirit or I wrote the Holy Ghost, since “ghost” is a fine, dark Saxon word, but “spirit” is a light Latin word. Then there is another reason. The reason is that I think that, of all languages, English is the most physical of all languages.

Buckley: The most what?

Borges: Physical. You can, for example, say “He loomed over.” You can’t very well say that in Spanish.

Buckley:Asomó?”

Borges: Well, no, no, they’re not exactly the same. And then you have, in English, you can do almost anything with verbs and prepositions. For example, to “laugh off,” to “dream away.” Those things can’t be said in Spanish. To “live down” something, to “live up to” something — you can’t say those things in Spanish. They can’t be said. Or really in any Romance language.

I’ve seen this interview before, but watched it again today after an email exchange with Kirk McElhearn. Quoting (with permission) from McElhearn’s email to me:

For many years, I worked as a French → English translator, and there is one key difference between the two languages. France is a Romance language, and English is a language with both Germanic and Romance (mainly French) influence. This means that English often has synonyms where other languages may not.

Using your example, “He leaped at the chance” and “He jumped at the opportunity”, both would be translated in French as “Il a sauté sur l’occasion.” Meaning that someone writing in French wouldn’t have the same range of words to choose from. It’s maybe not the best example, because both are clichés, but there are many examples of French words where English has both a Romance equivalent and a Germanic equivalent: pig and pork, sheep and mutton, beef and cow. Food words are just one example, but English also has many more verb choices than French, since it has a larger vocabulary coming from both influences.

English gleefully borrows from any and all other languages. McElhearn wonders whether English is thus more fingerprintable than other languages, because of its richer vocabulary of roughly equivalent synonyms, and its multitude of idioms. 


  1. However, this vein of pro-watermarking support from people opposed to AI in general has opened my eyes to the notion that Anthropic is throwing its support behind this in order to get people who despise AI off their backs. ↩︎


Apple TV Still Has No Start Date for ‘The Savant’ 

The Savant is a political thriller series starring Jessica Chastain that was supposed to debut a year ago. Apple “postponed” it, apparently out of fear of upsetting extremist right-wing nut jobs because the show is about an undercover investigator (Chastain) hunting down extremist right-wing nut jobs. Chastain was not happy about the show being delayed.

Last we heard about the show was back in April, when Marc Malkin reported this for Variety:

“Before it was like, ‘I don’t know if we’re going to see it,’ but now I can say, ‘We’re going to see it,’” Chastain told me exclusively on Saturday at the Breakthrough Prize ceremony in Santa Monica.

As for when, sources tell me that Apple is planning for a July release.

Given that it’s now the middle of August, I think that a July release is looking less and less likely by the day.

No Update Since Early July Regarding Siri AI Coming to the EU, Ever 

The Financial Times, back on July 1, with the transcontinental byline “Michael Acton in San Francisco and Barbara Moens in Brussels” (non-paywalled summaries from 9to5Mac and MacRumors):

Apple chief executive Tim Cook and EU tech chief Henna Virkkunen held “constructive” talks on Tuesday as the two sides aim to lower the temperature in a bitter dispute over the iPhone maker’s new “Siri AI.”

An EU spokesperson said the virtual meeting had involved a “constructive exchange on topics of common interest, on which the work continues”. The meeting included a discussion of how Apple can launch its reinvented Siri in Europe while avoiding millions of dollars in fines for violating the bloc’s flagship competition rules, according to two people familiar with the talks.

[9 paragraphs of explanatory backstory on the Siri AI/DMA standoff elided ...]

The dispute triggered a fierce public backlash against the commission, with European officials reporting hundreds of emails from consumers accusing Brussels of depriving Europeans of a new technology. One EU official said that a commission spokesperson had received a stream of abusive messages, including several death threats.

If there are actual kooks who made credible death threats, of course they should be investigated, identified, and arrested. But mentioning this in the context of the “fierce public backlash” feels like fishing for sympathy over a deeply unpopular policy of zero benefit to Europeans. That there are hundreds or thousands of complaints is no surprise. This policy sucks and is indefensible on practical grounds.

In November, Apple first proposed a technical fix to the EU it later dubbed a “Trusted System Agent” — a layer of software between a user’s device data and a third-party AI model. It would allow rival AI assistants to draw on personal information from the device without giving them full access to the data. However, Apple has yet to build the agent, and said it was looking for assurances from the EU before it starts.

A commission official said its contact with Apple on the idea was limited, and that it lacked a concrete proposal or details on how such an agent would work beyond the general concept. They said Apple “focused on obtaining a green light to delay the compliance”.

That’s the nut of it. Last winter Apple sent an entire team, including engineers, to Brussels to present a proposal for the TSA (maybe that’s another acronym that needs rethinking on the basis of prior art) to ask, basically, “If we build this, would you deem it DMA-compliant?” and the Commission’s response was basically, “Build it first and then we’ll tell you, after we get feedback from your competitors on what they think of it.” And Apple doesn’t want to spend up to two years building a complex system, exclusively for the EU, only to find out then whether it was all for naught.

As for how detailed Apple’s Trusted System Agent proposal was, we have a he-said/she-said dispute. Apple said at a press briefing at WWDC that it was quite detailed. An anonymous “commission official” here told the Financial Times “that it lacked a concrete proposal or details on how such an agent would work beyond the general concept”. Apple has more credibility here, if only because their statements claiming the proposal was detailed weren’t from anonymous sources. They were on the record.

In the meantime, we’ve seen what the European Commission is demanding of Google with Android regarding third-party LLMs. I can’t see Apple ever agreeing to such a system for iOS. We don’t know the details of Apple’s TSA proposal, but I’d be rather flabbergasted if it enabled the things the EC is demanding for third-party AI providers on Android, like unrestricted background execution.

The Information Profiles Cami Clark — Dario Amodei’s Wife, Ivanka Trump’s Friend, One-Time Would-Be Pornographer, and Anthropic’s ‘First Lady’ 

Cory Weinberg, Jemima McEvoy, Jessica E. Lessin, and Stephanie Palazzolo, writing for the paywalled-without-gift-links The Information:

As Amodei has hopscotched the globe to preach about the potential and risks of AI — from New Delhi to Davos to Sun Valley — Clark has almost always been near his side. Several people who know the couple describe Clark as Amodei’s emotional ballast, someone he has sought counsel from during the turbulence of Anthropic’s growth and clashes with Washington over the future of AI.

Clark may well be the most consequential person within Anthropic’s orbit who doesn’t have a formal role at the startup. You might even call her the first lady of Anthropic. [...]

And before Female Algorithm Technologies, Clark tried her hand at building a pornographic film company, Eddice, that planned to make movies with high-production values and female protagonists. She also hoped it would have online shopping capabilities. In 2011, she tried to raise money for Eddice from Jeffrey Epstein, who’d pleaded guilty to charges of soliciting a minor for prostitution three years earlier, according to emails made public as part of the Epstein files released by the Department of Justice earlier this year. She emailed Epstein the script for the first of a series of films they hoped to make (title: “American Girl in Paris”).

“We thought you and the ladies might enjoy,” Clark wrote. She then appended a winking warning: “A little nsfw,” an acronym for “not safe for work.” Epstein didn’t invest.

While enormous attention has fallen on AI and the people behind the companies developing it, Clark has largely flown under the radar.


Anthropic’s ‘Watermark’ Text Adulteration in Claude Is a Perversion of Writing

When I wrote this week about Anthropic’s announcement that all Claude models, worldwide, would soon begin “watermarking” everything they generate, including text, to comply with this EU regulation, we were left to speculate how this was going to work, because Anthropic offered not even a vague description of how it would work — despite the fact that the title of the announcement was, absurdly and insultingly, “How Claude Marks AI-Generated Content”.

My initial speculation was that maybe they’d hide invisible non-printing Unicode characters in the text. Just spitballing. Turns out that’s not what they’re going to do. What they’re going to do is apply a form of steganography, where the choice of words (or other token output) at inference time will leave fingerprints that can later, maybe, be detected probabilistically.

I initially guessed “invisible characters” not because I didn’t think of the semantic word-choice technique, but because I was a fool who took Anthropic at its word in their description of what they would do. Their original support document claims:

When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.

They say “imperceptible” and “doesn’t change the meaning, quality, or readability”. Their words. Not almost imperceptible. Not slightly changes the meaning, quality, or readability. That made sense to me, because that’s absolutely what I want — nay, demand — from any tools I use personally. It’s unacceptable for a tool to sacrifice an iota of clarity, coherence, meaning, quality, etc. for the purpose of embedding hidden clues within the text to suggest its provenance. That’s what I would and will demand. And Anthropic’s (original) support document unambiguously claims that’s what their system will enable. So if that were true, I couldn’t see what was left other than hiding invisible characters within the text.

My error was believing Anthropic that their system wouldn’t adulterate and corrupt the semantics of the text their models generate. That is in fact exactly what they plan to do. I should have my head examined for believing a single word of a document titled “How Claude Marks AI-Generated Content” that doesn’t explain, at all, how Claude marks (or will mark) AI-generated content.

How It’s Actually Going to Work

Yesterday, on an entirely different website than the original “How Claude marks AI-generated content” article (the one that didn’t explain anything at all about how it works), Anthropic published “How Claude’s Text Watermark Works”, which does actually explain in layman-accessible terms how it’s going to work. I will return to Anthropic’s new highly euphemistic and slightly misleading description below.

There’s a bunch of research on this topic, some of which I have also linked to below. But the very best description of the general idea behind the technique is an interactive essay by James Padolsey, “How AI Text Watermarking Works”. It’s a wonderfully cogent read, and the interactive elements splendidly illustrate the main concepts. A+ work. If you have any interest in this at all, I dare say you must read — and play with — Padolsey’s piece.

But here’s my stab at a layman’s high-level summary. If you toss a coin N times and note the results, you can determine with a degree of certainty whether the coin is fair or biased. LLMs are, in their popular incarnations, non-deterministic. Ask the same question of the same model and you often get at least slightly different answers. Maybe the same meaning, but different phrasing. At each decision point for generating the next token, the model makes a choice. With these semantic watermarking techniques, they make different choices for some tokens based on word lists that could be called “green” and “red”. At each decision point, they’re a little more likely to pick a word from the green list than the red list. That doesn’t mean they never choose words from the red list. Just that they’re less likely to than they would if the adulterated marking technique weren’t in place. (Same way that a crooked 51-49 coin will still land “wrong” side up 49 times out of 100 on average.)

Words or word phrases are sorted into the green and red lists deterministically on the fly, at each “next token” generation point. So sometimes a specific word will be on the green list, and other times it will be on the red list. Someone with the secret key can determine which list a word will be on at each token generation point (which is how the watermarking is detected); those without the secret key cannot. This means there will never be a list of words that Claude prefers or eschews.

With coin flipping, the higher N is — the more times you flip — the more confident you can be that the coin is fair or biased. So too with this semantic watermarking. The more words in the text, the more accurate the analysis will be that the text was generated by a specific AI model or not. With too few coin flips, you can’t achieve any confidence at all regarding a coin’s fairness. With too few words (or tokens), there’s no way to achieve any confidence whether a string of text was AI-generated or not.

Given a string of text to examine for signs of a specific watermarking system, if there are more words tagged as green and fewer tagged as red than would otherwise be expected, the text can be flagged — with some degree of confidence — as having been generated, or merely modified, by the AI system that applies the specific secret-key watermarking system. The amount of confidence in the determination will obviously vary, significantly, based on the size of the text string and randomized weights given to words on the green and red lists. But only Anthropic will be able to determine if text was seemingly generated by Claude, and Anthropic will only be able to detect the watermarks that are applied by Claude. Claude can’t detect the hidden watermark signals generated by, say, Gemini, and Gemini can’t detect the hidden watermark signals created by Claude, because each implementation is predicated on secret keys held only by the LLM provider.

Objections to the Technical Premise

One of my fundamental problems with this is that no two synonyms carry the exact same meaning. “He leaped at the chance” and “He jumped at the opportunity” are very similar sentences expressing the same general sentiment, but they are not the same. The exact words we choose when writing matter. I want any LLM I use to choose the very best, most precise words at every single decision point. An obvious constraint that I accept is time and computation. Within the constraint of executing inference quickly, and at a certain cost per token, I want the best words. This constraint matches human writing. I could surely write a better column by taking longer to write it. I write with a sense of how much care I should put into every word and punctuation choice I make. I take more time with certain paragraphs, sentences, or even individual word choices when my gut feeling says I should.

In other words, these are necessary trade-offs. These factors are all in my interest: speed, cost, quality. Ideally I would like perfect writing, at instantaneous generation speed, at zero cost. None of those things are possible. Computation is not free of charge (and cloud-based LLM inference with leading models is actually expensive). Inference is not instantaneous. And great writing, whether natural or artificial, can only approach perfection.

The idea that anything other than my needs should factor into the generation of text for me is patently offensive.

This isn’t just about text one might generate with the intention of passing it off as their own natural work. This isn’t even about LLM proofreading of work written by hand. Anthropic is saying that all new Claude models are going to adulterate every single bit of text longer than 200 tokens (~150 words) they generate, including everything it presents to its users to read. So even in a private conversation between a user and Claude, which will never be read by anyone other than the user, Claude will begin making word choices in the name of marking its output in statistically predictable ways rather than maximizing clarity and precision.

Even today’s so-called frontier models are already decidedly lacking in lucidity. Claude, ChatGPT, Grok, et al. are “better writers” than most humans and produce better prose than the median human. But: no shit. Most people are terrible writers. The “average person” is pretty stupid and half of all people are stupider than that. And there are many smart, interesting people who are miserable writers. So as impressive as LLMs are, the bar is low. The best writing I see come out of these models is worse than anything I would choose to read for pleasure. And now Anthropic is saying they’re going to make it worse, on purpose, for purposes that do not benefit me in any way? Even if only slightly worse?

Get fucked.

Objections to the EU Regulation

Speaking of objections, the relevant EU regulation motivating all of this, “Code of Practice on Transparency of AI-Generated Content”, is red-tape nanny-state pipe-dream nonsense. Here’s Ben Thompson’s summary from a paywalled Stratechery update this week:

  • The regulation applies to text longer than 200 tokens.
  • The provider must mandate in their terms-of-service that users not remove the watermarking.
  • The solution should be robust in terms of evading “typical processing solutions” like screen shots, scanning and OCR, copy-and-pasting, translations, etc.

Taken literally, compliant LLM terms of service must forbid users from rephrasing the output from models that comply with this regulation, because the word choices are the marks. But it’s not the European Union that is trying to impose their absurd, impractical, witch-hunt-fueling regulation on the entire world. That falls on Anthropic.

Complying with this, particularly with regard to text, is only going to create problems for honest users. Dishonest users attempting to pass off AI-generated text as their own writing (students, employees, whoever) will simply circumvent detection through non-compliant AI paraphrasing tools.

James Padolsey — whose interactive visual explanation of how these schemes work I linked to above — explains this in a post titled “Anthropic’s Weak Watermarks Appease a Weak Law” (which, if it rings a bell, I linked to in a standalone post earlier today):

The same thought that led to this law could have applied to calculators at the time of their inception, had their outputs revealed themselves through artefacts. Thankfully, a sum borne of the brain is treated no differently from one produced by a calculator. Likewise with spellcheckers. To make assistance suspect only once the tool becomes capable enough to compose a whole sentence is not a principled boundary. It is a moral premium placed on difficulty itself.

Anthropic has nevertheless chosen a blanket, model-level implementation that appears broader than the law’s minimum requirement. That may be convenient compliance engineering, but it discards distinctions the law expressly attempted to preserve. The result is a signal broad enough to implicate harmless and assistive use, yet fragile enough to be removed by a motivated person through substantial recomposition. It risks concentrating suspicion on ordinary and assistive users while remaining weakest against deliberate deception.

Padolsey is the creator of Declaude, a delightfully simple web app that allows you to “Paste in AI-flavored text and get the same content back as plain prose”. Declaude’s original purpose is cleaning the saccharine Claude personality stink from text (whether it was created by Claude or any other LLM), but, if Anthropic persists in its stated plan to begin adulterating all text Claude generates, Declaude will also serve as a copy-paste single-extra-step way to eliminates those marks. Declaude is interesting and useful already, but it exemplifies how ill-considered and futile this EU regulation is when it comes to prose.

Google SynthID

Google has a watermarking system in place that they call SynthID, which they apply to AI-generated images, video, audio, and text. I’m concerned in this article only with text. With multimedia, embedded watermarks can be metadata within files, and truly not affect the experiential quality of the work when viewed or listened to. With text, we are talking about the actual words that are chosen. From the “AI-generated text” section of Google DeepMind’s own description of SynthID:

We’ve expanded SynthID to watermarking and identifying text generated by the Gemini app and web experience. Large language models generate text one word (token) at a time. Each word is assigned a probability score, based on how likely it is to be generated next. So for a sentence like “My favorite tropical fruits are mango and…”, the word “bananas” would have a higher probability score than the word “airplanes”. SynthID adjusts these probability scores to generate a watermark. It’s not noticeable to the human eye, and doesn’t affect the quality of the output.

In a group chat, a friend of mine quoted the above, and I responded that if a chatbot wrote “My favorite tropical fruits are mango and airplanes”, I’m pretty sure I’d fucking notice. Another friend then responded with this:

AI-generated image of an airplane carved out of a pineapple or something, on a tropical beach.

Days later, that still cracks me up.

But Google’s absurd description puts the lie to their own claim that it isn’t noticeable, and it serves to show just how little regard the people behind these generated-text fingerprinting schemes have for the actual craft of writing. Of course bananas has a higher probability score than airplanes, because airplanes aren’t fruit. But what about pineapple? Should the sentence complete to “mango and bananas” or “mango and pineapple”? That’s a good question, and the only acceptable answer for why an LLM should choose bananas instead of pineapple (or coconut, or guava, or papaya...) is that it has determined that it’s the best fit for the intended meaning, tone, and sentiment of the text. Not because bananas is on the watermarking “green” list and pineapple is on the “red” list, even though pineapple might be the better fit. Google’s own supposedly jocular description of how SynthID works in fact captures how the scheme perverts the text it generates.

They’re saying you won’t notice because if it only chooses bananas over pineapple for these fingerprinting purposes, well, they’re both tropical fruits and who cares. But it’s utter nonsense that the difference is “not noticeable to the human eye”. The semantic difference between banana and pineapple is just as noticeable to the human eye as the taste of the two are to the human tongue.

If it did produce “My favorite tropical fruits are mango and airplanes”, it’d be incredibly stupid, but it wouldn’t be offensive because we’d all recognize that something completely off-key happened. What’s offensive is that with a system like SynthId in place, where the fingerprinting decisions are motivated by a secret key, we have no idea whether it completed to “mango and bananas” because bananas was determined to be the best next token, or because bananas is in the “green” bucket of words. It calls every single word choice into question.

Here’s a paper published in Nature where Google’s team behind SynthID published their work, after putting it into production with Gemini (née Bard):

We analysed approximately 20 million watermarked and unwatermarked responses and computed the thumbs-up and thumbs-down rates (both as a fraction of the total number of thumbs-up and thumbs-down feedback received). We found that the thumbs-up rate for the two models differed by 0.01% (with the watermarked model being higher); and the thumbs-down rate differed by 0.02% (with the watermarked model being lower). We found both of these differences to be statistically insignificant, and well within the 95% confidence intervals.

From this experiment, we conclude that over a wide variety of real chatbot interactions, the difference in response quality and utility, as judged by humans, is negligible. Subsequently, non-distortionary SynthID-Text has been productionized and is currently watermarking responses in Gemini and Gemini Advanced. To the best of our knowledge, this evaluation represents the first systematic watermarking investigation of its kind within a large-scale production system.

To this I say:

  • Gemini/Bard’s thumbs-up/thumbs-down buttons are not a good experiment for evaluating the effect on quality. If a chatbot tells me “My favorite tropical fruits are mango and bananas” instead of “mango and pineapple”, I’m not going to give the response a thumbs down because of the fruit it chose. I’d give it a thumbs down if it said “airplanes”, yes, but that’s a strawman. (The paper in Nature even uses “My favourite tropical fruit is ...” as an illustration, but in the paper, the only four next tokens considered are, in order of probability distribution, mango, lychee, papaya, and durian. No airplanes. And, conveniently, in the paper’s example, the “winner” of the watermarking “tournament” just happens to be mango, the one that would have been selected as the best if the watermarking weren’t in place.)

  • A “difference in response quality and utility, as judged by humans” that is “negligible” does not mean imperceptible. What they really mean is that it’s only slightly worse and that everyone is either too stupid to notice or too indifferent to care.

  • It’s widely considered that Gemini is behind ChatGPT and Claude in quality. Perhaps the fact that they’ve put SynthID-text into production is one of many reasons why. I personally agree that Gemini’s prose is inferior. Maybe the use of SynthID has nothing to do with the fact that I, along with the general public consensus, consider Gemini to be a second-rate chatbot — but in that case, maybe it’s the fact that Gemini is a second-rate chatbot that makes the difference “negligible” when Google started mixing in SynthID-motivated tokens in its results. It’s a lot more likely that your restaurant customers won’t notice that you replaced your regular coffee with Folgers Crystals if your regular coffee is second-rate to start with.

Anthropic

Now, finally, back to Anthropic’s new “How Claude’s Text Watermark Works”, published yesterday. I have some comments.

To summarize:

  • We use a method of watermarking that does not have any practical impact on the quality or content of Claude’s outputs;

  • The difference between watermarked and un-watermarked text will not be distinguishable to readers;

Translation: Specific words do not matter and we don’t think anyone reads anything closely.

  • Nothing is added to the text and there are no hidden characters;

This would have been worth clarifying at the outset.

  • Watermarking won’t be specific to Claude. As of August 2, the EU requires AI providers serving its market to mark AI-generated content. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks.

No other AI provider has stated that they will apply such marking, adulterating all generated text, outside the EU.

Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses — the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number.

Arguing that grey vs. overcast “doesn’t matter much to the reader” is the crux of my argument that this entire endeavor is a perverse adulteration of what it means to write — or to read. That it’s subtle in some ways makes it more perverse, because it’s sneaky.

In internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text. In the SynthID-Text paper, which introduced the technique we use, Google DeepMind tested this impact by serving a model that used watermarking to a portion of their Gemini traffic and comparing thumbs-up and thumbs-down ratings. They found no statistically significant differences from the unwatermarked model. And in a controlled study, human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality.

See above for my argument that this thumbs-up/thumbs-down data is absolutely worthless in evaluating whether the SynthID-style word-bias watermarking makes text worse. By definition it must make text worse, unless the underlying LLM model’s scoring is wrong, because the nature of the watermarking algorithm requires it to sometimes increase the probability of selecting a worse word choice and decrease the probability of selecting the model’s best choice. It’s only a question of how much worse. What Google’s thumb-counting data shows is only that it isn’t so much worse as to make Gemini users click the thumbs-down button.

Watermarking doesn’t change the meaning or experience for the person reading it, but if you wanted to check after the fact whether the text was likely generated by Claude, the watermark allows you to do so.

No, it does not. Because the entire scheme is tied to secret keys held only by the AI provider, it only allows Anthropic, not “you”, to check anything.

When Claude proofreads text written by a person, what it gives back has generally only been lightly edited; because nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to. Depending on the length of the text and how heavily Claude has edited it, those changes might not be enough to make Claude’s involvement detectable. The more Claude writes, the more decisions it has to make, and the more space there is for a watermark.

Translation: No one can ever again use Claude for proofreading their own prose unless they’re willing to risk that the whole thing might be flagged as having been generated by Claude.

For example, once the model has written “2 + 2 =”, there is a very clear best choice for the next token (if the model is completing the sum, there isn’t an answer that’s equally as good as “4”; if it’s talking about George Orwell’s Nineteen Eighty-Four, there isn’t an answer that’s equally as good as “5”). The “nudge” of the watermark wouldn’t be applied here. For the same reason, code — which in very many cases has to be exact — has generally less watermarking than some other forms of text.

Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code. But by definition, it will have a negligible effect on the actual code produced.

Translation: We value precision in programming code; we do not in prose.

And it is exceedingly rich to cite George Orwell’s Nineteen Eighty-Four, approvingly, in the context of justifying a text adulteration scheme premised on the notion that specific words do not matter. I mean what the actual fuck? Orwell!

Lastly, as to why they’re doing this:

We’re implementing watermarking to comply with the EU AI Act. Anthropic, along with several other major AI model providers and around 190 total signatories, signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026. This requires AI system providers to use methods of “marking” AI-generated text. We’re applying watermarking globally at launch because we don’t yet have a durable way to scope it by region.

This, from a company that the Financial Times just reported is weeks away from an IPO with an intended valuation of $2 trillion, which would make it one of the 10 highest-valued companies in the world — as of today, placing it at #7, between TSMC ($2.2T) and Broadcom ($1.9T).

This leaves us to believe that one of the following must be true:

  • It’s perfectly reasonable that a technology company valued on par with Amazon and TSMC is technically incapable of complying with an EU regional law only within the EU itself.1 Not a cause for concern at all.

  • Anthropic is in over their heads, wields shockingly little control over their own tech stack, and their imminent IPO is likely to be remembered only as a new high-water mark in the manic global AI bubble.

Also, what happens if another major global market makes it unlawful for AI to secretly watermark generated text?

OpenAI

From an OpenAI support document titled “Provenance Signals (Content Credentials, SynthID) in OpenAI-Generated Content”:

Consistent with our commitments under the European Commission’s Code of Practice on Transparency of AI-generated content, our goal is to expand provenance signals to all modalities including text, so customers and developers have clear ways to meet their own transparency obligations as standards and tooling continue to mature.

There’s a lot of wiggle room in this brief statement, and it could just as well mean that OpenAI models will only adulterate text with fingerprint markers when users or developers ask for it. Or that it will only be mandatory for users in the EU. If I were at OpenAI I’d go hard on this and publicly say that ChatGPT will never watermark text it generates unless you ask it to, and that if you want tools that secretly work behind your back without telling you how they work to flag your words in ways you can’t see, go ahead and use Claude.

Further Reading

Three papers on ArXiv:

I will admit that while I’m profoundly offended by the idea of personally using tools that attempt to leave such watermarks in text they produce or touch, the mathematics behind it are fascinating.

Michael Lopp, at Rands in Repose, “RIP Claude”:

As a human who has had to wrangle with EU regulations in the past, I am abundantly clear what’s involved in the laborious bureaucratic process. I can guess what threats Anthropic is facing. However, this is a tone-deaf, clumsy, and alarming opening salvo in their watermark strategy. [...]

My writing is my work, and Anthropic’s current strategy is aggressively writer-hostile.

Jeff Gamet, “Anthropic’s Claude Watermark Is Akin to an AI Poison Pill”:

To be clear, the watermarking is embedded in pretty much any text Claude touches. Along with text Claude generates, it also applies to text it processes, such as proofreading and summarizing. I expect we’ll see too many inaccurate accusations of using Claude to write documents where the content was human-written, but AI-proofread.

The watermarking sticks with documents through copy-and-paste, too. Imagine copying text from a blog post or email only to have what you wrote tagged as potentially AI-generated. In fact, that could very well happen with this post. I personally write all of my content without AI tools, but I copied the quote at the top of this piece directly from Anthropic’s website. Does that mean what I wrote here will show as AI-generated? If they used their own models to generate or edit what I quoted, then the answer is very likely “yes.”

One of the papers published at ArXiv I cited above claims that such watermarking even persists when an article of text originally generated in English is translated into German.

Secrets are the poison here. When only Anthropic holds the secret keys that both produce the watermarking and perform the probabilistic detection of those marks, we’re all left to wonder. To wonder if what we’re reading is secretly watermarked, what we’re quoting is secretly watermarked, and whether what we ourselves are writing will be unjustly accused of being AI-generated based on secrets we don’t know and can’t see. Poisonous is exactly the right word.

Or should I say toxic? Or airplanes


‘Anthropic’s Weak Watermarks Appease a Weak Law’ 

James Padolsey, on the Claude-text-watermarking-to-comply-with-an-EU-regulation imbroglio:

The same thought that led to this law could have applied to calculators at the time of their inception, had their outputs revealed themselves through artefacts. Thankfully, a sum borne of the brain is treated no differently from one produced by a calculator. Likewise with spellcheckers. To make assistance suspect only once the tool becomes capable enough to compose a whole sentence is not a principled boundary. It is a moral premium placed on difficulty itself.

Anthropic has nevertheless chosen a blanket, model-level implementation that appears broader than the law’s minimum requirement. That may be convenient compliance engineering, but it discards distinctions the law expressly attempted to preserve. The result is a signal broad enough to implicate harmless and assistive use, yet fragile enough to be removed by a motivated person through substantial recomposition. It risks concentrating suspicion on ordinary and assistive users while remaining weakest against deliberate deception.

Padolsey is the creator of Declaude, a delightfully simple web app that allows you to “Paste in AI-flavored text and get the same content back as plain prose”. Declaude’s original purpose is cleaning the cutesy Claude personality stink from text (whether it was created by Claude or any other LLM), but, if Anthropic persists in its stated plan to begin adulterating all text Claude generates, Declaude will also serve as a copy-paste single-extra-step way to eliminates those marks. Declaude is interesting and useful already, but goes to show how ill-considered this EU regulation is.

Padolsey also wrote and programmed a splendid interactive essay that explains and illustrated how the “watermarking” scheme Anthropic is adopting (and Google Gemini has already adopted) works.

Trump Administration ‘Not in Favor’ of Apple Using Chinese RAM 

The Wall Street Journal (gift link):

To help alleviate the supply crunch, Apple is looking to Chinese manufacturers.

“The Trump administration is not in favor of that,” Lutnick said in an interview after touring an Apple manufacturing facility in Houston. There have to be “other solutions to the memory issue, but it’s not great American companies using Chinese memory.”

Asked if he has relayed that message to Apple, Lutnick said “plainly.”

But:

U.S. government rules require American companies to secure a license before sharing product information with CXMT and YMTC. The Chinese chip makers would need such information if they were to make customized chips for Apple. But Apple is free to buy off-the-shelf parts from the companies, and to negotiate prices with them.

Apple’s chief operating officer, Sabih Khan, declined to confirm that Apple was testing Chinese memory chips in an interview Thursday. But he said that given the magnitude of the supply shortage, “we have to look at all options,” including working with existing suppliers to increase U.S. production.

XCancel — An Unofficial Twitter/X Mirror 

XCancel:

XCancel is an instance of Nitter.

Nitter is a free and open source alternative Twitter front-end focused on privacy and performance. The source is available on GitHub at https://github.com/zedeus/nitter [...]

Using an instance of Nitter (hosted on a VPS for example), you can browse Twitter without JavaScript while retaining your privacy. In addition to respecting your privacy, Nitter is on average around 15 times lighter than Twitter, and in most cases serves pages faster (eg. timelines load 2-4× faster).

I personally don’t care for Elon Musk’s X-rebranded Twitter. I also realize that some of you outright despise it, and, worse, that under Musk, they make it difficult to view content if you’re not signed into an account or if you have JavaScript disabled. So, sometimes, when I link to posts or threads on Twitter/X, I include a link to the same content at XCancel.

XCancel isn’t perfect, alas. Just this week I linked to a tweet from Google Design that contained an animation of an Android app that they were inexplicably proud to show off. The XCancel cache of that tweet only showed it as animated to some people. Others just saw a static image. I don’t know why. That’s XCancel’s problem, not mine.

Also, it’s tiresome and repetitive to include an extra link to XCancel every time I link to a tweet on Twitter/X. If you would prefer to view x.com links on xcancel.com, you should automate the redirection. If you use Safari, XCancel Redirect is a free extension that claims to do just what you think it does. (I haven’t tried it.) Or you can use a general-purpose redirect extension. RedirectWeb is a good one that I can recommend. Also, Jeff Johnson’s excellent StopTheMadness — an extension I rely on for multiple purposes and frequently recommend. With StopTheMadness, you can redirect all x.com URLs to xcancel.com with the following rules (you need the leading and trailing slashes in the pattern to tell STM that it’s a regex pattern, not a plain text string):

Pattern: /https?://(www[.])?x[.]com/
Replacement: https://xcancel.com
Enabled on platforms: All

I understand not wanting to visit Twitter/X. But if you feel that way, I do sympathize, but that’s your decision, and you should use tools to redirect Twitter/X links automatically. And if you’re wondering why I don’t just refuse to link to anything on Twitter/X, that just isn’t practical. I wish it weren’t so, but many people, including senior Apple executives, and many organizations post interesting content exclusively to Twitter/X. If something posted to Twitter/X is also available elsewhere, I link to the elsewhere version. But if it’s exclusively available on Twitter/X and interesting, I link to the original.

My linking to original content on Twitter/X does not perpetuate Twitter/X’s relative popularity. Original content appearing exclusively on Twitter/X does.

Drata 

My thanks to Drata for sponsoring last week at DF. Their message is short and sweet: Leverage autonomous AI agents to automate compliance, manage internal and third-party risk, and continuously prove your security posture.


You Don’t Need to Worry About Scratching Your iPhone Camera Lenses

Following up on my post yesterday about modern iPhones and scratch resistance, and my personal habits of (a) almost never using an iPhone case, and (b) never setting the iPhone face down except on soft (cloth) surfaces.

I should have anticipated this, because I’ve had this conversation with real-world normies repeatedly in recent years, but I got a bunch of questions from people who say they’re wary of ever putting their iPhone down on its back because they’re worried about scratching the camera lenses. That’s not a silly thing to worry about. It looks like those lenses are glass; glass scratches; and it sure seems like a scratched camera lens might forever ruin all the photos and videos you take with the camera.

There are a couple misconceptions here though. First, the glassy flat circles exposed on the back of your phone aren’t the camera lenses. Those are covers over the lenses, which are smaller, spherical (convex, not flat), and recessed. You can look through the covers and see the actual spherical lenses inside. The exposed lens covers are made of sapphire, not glass, and are thus incredibly scratch-resistant. They’re by far the most scratch-resistant parts of the phone. I’ve never once found even a tiny scratch on any of my iPhone camera lens covers, and I’ve never taken any particular care to avoid scratching them. I mean, I don’t drag the lenses face-down on surfaces. I’m not trying to scratch them. But I set my phones down on hard surfaces lenses-down all the time and they never seem to pick up even fine scratches.

Zack “JerryRigEverything” Nelson is the guy on YouTube who makes videos where he scratches the hell out of products to see how durable they are. (And he bends them, burns them, and abuses them in other gruesomely creative ways.) In his iPhone 17 Pro video, starting around the 4:10 mark, he tries scratching the sapphire lens covers with a sharp razor blade. No effect. Sapphire is far more likely to shatter than scratch, because it’s so hard. (You may recall that a decade ago Apple pursued using sapphire for the displays of iPhones but it didn’t work out.) Don’t try scratching your lenses with a diamond hardness pick and you’ll be fine.

Also, believe it or not, a fine scratch on the lens cover will almost certainly not affect image quality at all. Try sticking a strand of hair (which is probably much thicker than a typical scratch) to the surface of your iPhone 1× main camera lens. Take a picture. Clean the hair off the lens. Retake the same picture. You almost certainly won’t see any difference at all. Here’s a great video from DPReview back in 2020 showing just how much dust or scratching you need to put on the outside of a lens to degrade image quality. It’s counterintuitive but the surface of the lens is not where light is focused — the sensor, inside the camera, is. 


Google’s ‘Material 3’ Design Write-Up Is 93.3 Percent Embarrassing 

This page from Google Design on their “Material 3” UI language came to my attention after my snarky post about the ungainly new to-do app they bizarrely bragged about on Twitter/X this week. I don’t think this “Material 3” page is new — I think it’s a few years old — but I’d never seen it before.

First, it’s crazy that (in desktop browsers with a mouse cursor) they change the I-beam cursor for text selection to ... a circle. I guess it looks kind of fun but the whole point of the I-beam cursor is to enable precise character selection. They changed that to something that makes precise selection as difficult as possible. It’d be better to just use an arrow pointer than to replace the I-beam with a goddamn circle. This, from the UI design team.

Second, this paragraph made me honestly wonder if I was reading a spoof, a years-old unfunny April Fool’s “prank”. But as far as I can tell this is totally straight, not a prank:

These factors can be quantified in users’ responses to new M3 Expressive designs. We found a 32% increase in subculture perception, which indicates that expressive design makes a brand feel more relevant and “in-the-know.” We also saw a 34% boost in modernity, making a brand feel fresh and forward-thinking. On top of that, there was a 30% jump in rebelliousness, suggesting that expressive design positions a brand as bold, innovative, and willing to break from convention.

Nothing says rebellious and “in-the-know” like assigning precise percentages to “rebelliousness”, “modernity”, and “subculture perception”.

Update: Believe it or not there is video of the Google Design team putting this research together. Worth a watch.

Ceramic Shield 2 Is the Real Deal 

Philip Michaels, writing last September for Tom’s Guide:

iPhone 17 torture test videos done by JerryRigEverything indicate that Ceramic Shield 2 certainly resists scratching, with scratch testing leaving only light scratches at level 7 on the Mohs scale of hardness. Scratches typically show up on glass at levels 5 or 6 on that scale.

“Ceramic Shield 2 is indeed the best we’ve ever seen,” JerryRigEverything remarks in the iPhone 17 Pro testing video.

That backs up Apple’s own Ceramic Shield 2 video, in which a mineral tip can be seen rubbing against a Ceramic Shield 2-coated display. There’s residue left on the screen, but it’s material from that tip, as it wipes away fairly easily.

I mentioned yesterday that I try never to place my iPhones face down, to avoid scratches, and that my nearly year-old iPhone 17 Pro seemingly has not one visible scratch on the front glass. Not even a single micro abrasion I can see, even when tilting it with the display off to hunt for them. Surely Apple’s Ceramic Shield 2, which is intended to provide best-ever scratch resistance, is a factor too — and likely the biggest factor. Apple at its best.

Google Design Pisses Its Pants on Twitter/X 

A few years ago I’d have looked at this post and maybe leaned toward the idea that a precocious 8th grader somewhere hacked into the @GoogleDesign Twitter account and tried to pass off their little to-do app as having come from Google’s design team. But this is apparently real. I almost hope it’s AI slop and that there aren’t any human designers there who think anything in this app has appropriate proportions or is aesthetically pleasing.

(Here’s an XCancel link for the x.com averse. You can just change any x.com/* URL to xcancel.com/*. But, alas, the XCancel mirror of the original tweet loses the animation showing the UI in motion — for some, but not all, users.)

Joanna Stern on the Pixel 11 ‘HiLight’ Notification Light 

Joanna Stern, writing at The New Things (gift link):

I used to love the blinking notification light on my BlackBerry, and later my Droid 2. It was a simple way to know I had a message without actually looking at my messages. Then BlackBerry let you customize the color, and it was a rainbow dream.

Google’s HiLight takes it a step further by letting you assign different colors to VIP contacts. So when your phone is face down, you can tell who’s trying to reach you without picking it up. It looks cool. The big bummer? At launch, it only works for phone calls.

Google told me it’s “continuing to invest in this technology” and that messaging notifications are coming. Still, it’s odd they aren’t there at launch, especially since Google found plenty of colors for Gemini. The ring changes hue depending on whether the AI is listening, thinking or responding.

I have no interest in this because a few years ago I stopped ever putting my iPhone on any surface face down. I don’t use a case, so setting the phone face down is a good way to pick up scratches. I stopped doing that, ever, and now I’ve got a nearly year-old iPhone 17 Pro that seemingly doesn’t have even a micro abrasion on the display glass.

And it seems downright goofy that it’s launching with support only for phone calls. Google can’t even launch a blinking light right on these Pixel phones.

Google Introduces ‘Camera Looks’ With Pixel 11 Phones 

David Imel, The Verge:

But in our current moment, the photos people are drawn to are not flawless — and that’s created a real problem for the people making smartphone cameras. “The gap between what two random people want from their camera is growing dramatically,” says Isaac Reynolds, who leads the Pixel camera team at Google. Some people want a perfectly optimized photo, Reynolds says. “But there’s a growing number of people who want something that they feel is more authentic or traditional, by their definition.”

That’s what Reynolds and the Pixel team are setting out to solve with Camera Looks, a total rethink of how the camera captures and styles an image on the Pixel 11 series. The phones include a new suite of processing tools aimed at giving people much more control over the camera’s output. At first glance, they look a lot like filters, with names like Black Tie, Classic, and most strikingly, Digi, in reference to the growing trend of buying cheap digital cameras from the early aughts. But unlike a filter, which generally applies editing tweaks after the processing happens, Camera Looks is changing how image data is processed at the sensor level.

This sounds a lot like Apple’s “Photographic Styles” in the iPhone 13 and later, but with the addition of some control over grain and sharpness. I love grain, and the thing I like least about Apple’s built-in photo processing is the way it attempts to eliminate all noise. I like digital cameras that embrace noise and try to make it pleasing, rather than over-processing to try to pretend these tiny sensors aren’t inherently noisy. So I say kudos to Google for loosening up on “everyone gets the over-sharpened, over-smoothed look”.

But, that said, what’s made me happy over the last year is shooting almost all my photos with third-party apps — in particular, Halide, Not Boring Camera, and Analogue — that give me low-level control over processing through LUTs (or in Halide’s parlance, “looks”). These apps give me the looks I want just through point-and-shoot photography. No need to process images after shooting. But because they (optionally) shoot RAW, I can reprocess in post without losing any original sensor data. I can shoot in black-and-white and subsequently change my mind and reprocess with a color look later. These apps enable honest film-look photography without looking at all gimmicky, and without baking any sort of filtering into the image data in my camera roll. It’s the best of both worlds. It’s also way, way fussier than the built-in Camera app should be.

Anyway, all this Pixel announcement news today has me thinking that these days, I never hear much about Pixel cameras or image processing. In the early years of the Pixel lineup, there was a common view that Pixel phones were the best camera phones in the world. I never believed that, but I did think they were an interesting alternative, taking different but equally valid philosophical choices than Apple. Apple has always insisted on instant processing. You snap the shutter and the iPhone saves your image instantly. Pixel cameras often undertook seconds-long processing. That enabled features the iPhone couldn’t offer.

I think it’s pretty easy to pinpoint when things changed for Pixel — when Marc Levoy left Google for Adobe in 2020 (where he now leads the team making the very intriguing Indigo camera app, that is exclusively available for iPhone and which offers a slew of innovative computational photography features).

When Levoy was leading Google’s camera photography team, there was a credible argument that Pixel phones were the best still camera phones in the world. (They were never even close on video.) Ever since, I never hear about them except in the middle of August each year, when Google announces new Pixel hardware. See you in 12 months.

TechCrunch on Google’s Pixel 11 Lineup 

Ivan Mehta, TechCrunch:

With this year’s Pixel 11 series launch, Google is thinking more agentic AI to complete your tasks. With Gemini, U.S.-based users will be able to order groceries, book rides, or get coffee, for instance. Plus, Gemini can call businesses on users’ behalf for table reservations or appointments. Google said that users can take over or stop tasks at any point in time and also review transcripts for AI-operated calls.

Over the next few weeks, a new slate of connected apps will be able to work with Gemini to get things done, including Granola, Otter.ai, Wix, Fever, Get Your Guide, Localiza, OpenTable (U.K.), Ticketmaster, iHeartRadio, Pandora, Angi, Thumbtack, and Zocdoc.

I’ve never heard of most of those apps, or, if I have heard of them, I thought they went defunct years ago. The one app amongst those I do use, OpenTable, apparently only supports this Gemini integration in the U.K.?

Another knockout year for the Pixel phones and Android.

Hands-On With Google Pixel 11 Pro Fold 

Sam Rutherford, writing for Engadget:

Granted, the P11 Pro Fold still includes IP68 dust and water resistance, which is better than what you get on the Z Fold 8 line (IP48). But after making an absolute tank of a foldable phone with last year’s model, I said Google really needed to cut weight and thickness this generation, and it just hasn’t. Instead, it seems like Google has leaned even more into the Pro Fold’s sturdiness with a new glass fibre material for its rear panel that the company says is nearly impossible to crack. This is great in theory, but obviously I wasn’t in a position to smash one of Google’s demo units in order to test that claim.

I have never seen anyone, anywhere, with a folding Pixel phone. Pretty sure I won’t this year either.

I played with a Pixel 10 Fold or whatever its name was last year at the Google Store on Newbury street in Boston, which store felt like a small library or museum. The Pixel Fold display was so plasticky that I had a moment where I wondered if the display units were fake props, like at Ikea. But no, it was the real thing. And it cost $1,800 to start.

Google’s Pixel Watch 5 

Victoria Song, The Verge:

The $399 Google Pixel Watch 5 isn’t about the hardware. Sure, there’s a new satin pyrite case finish, a few new strap colors, and a Steph Curry Special Edition. Under the hood, there’s a slightly faster Qualcomm processor and an itty-bitty battery bump. There’s a $50 price hike from last year, too, because the Pixel Watch 5 isn’t immune to RAMageddon — none of us are. Otherwise, no one would blame you for looking at this watch and thinking absolutely nothing’s changed. That’s because the big updates this year are all software-based.

I have never seen anyone in real life wearing a Pixel Watch. (I’ve only seen a handful of people using Pixel phones, and every single one of them was at a tech event of some sort.)

You can also use AI to generate watchfaces. I created a watchface featuring “fat cats and an elegant font.” It didn’t quite do that. The fat cats weren’t characters, so much as pleasantly plump, cat-shaped clock numbers. It’s a gimmicky way to shoehorn the Nano Banana model onto the watch… but if that’s your thing, it works.

Well I’m sure now I’ll see plenty of people wearing Pixel Watches.

Amazon Is Spiting Customers With Unhelpful Order Confirmation Emails 

Mia Sato, The Verge:

Earlier this summer, Amazon customers began noticing that emails related to their online orders looked sparse: Order confirmation emails didn’t name specific items anymore, and instead listed only item categories.

“Your Beauty item is confirmed!” an email about my retainer cleaning tablets read. Shoppers have posted other iterations of the redacted emails as well: “Ordered: 1 Hardware item,” “Your Drugstore, Shoes, and other items are here!” and “1 Nutrition & Wellness, 1 Wireless Accessories,” for example. The emails have clip art-style illustrations of general product categories, and a shopper has to exit their email and go to Amazon to see what item it’s actually referring to.

Complaints about the vague emails appear to have started in July; orders placed as recently as June contain thumbnails and names of the exact item purchased. [...]

“As customers shop with us more frequently, including on their phones, and use the ‘Your Orders’ page in the Amazon app to get real-time, consolidated order details and delivery status, we’ve simplified several order-related emails to direct customers to our app and website for the latest information on their orders,” spokesperson Maxine Tagay said in an email. “This also reduces customer information shared outside the Amazon app and website to further improve customer privacy.”

That statement from Amazon is almost honest, if you squint at it and ignore the bit about it having anything to do with “simplification”. If you order a Torx screwdriver, what’s simple is getting a confirmation email with the exact name, brand, and photo of the item you ordered, not “a hardware item”. The honest part is “reduces customer information shared outside the Amazon app”.

The fact is, a lot of people use apps like Gmail for email. In fact, a lot of people use Gmail in particular. There’s a good chance you do. When you use Gmail, Google learns everything in your email. Amazon sees this as a competitive problem for them, not a privacy problem for users. Everyone who uses Gmail basically knows that Google “knows” what’s in their email, and they’re voluntarily agreeing to that. Amazon just doesn’t like that any of its competitors — perhaps especially Google — can glean so much information about what people are buying from Amazon. So Amazon, in an act of competitive spite that corrupts its relationship with its own customers, is limiting all useful purchase information to channels that it controls — its website and its own apps.

If customers relied solely on the “Your Orders” page at Amazon, they wouldn’t need to send any emails at all. The obvious fact is, people go to Amazon when they want to buy something. They don’t come back until they want to buy something else. They enjoy getting confirmation and shipping notifications in email because they check their email regularly.

App Store Scam of the Week: ‘TabControl Extension’ for Safari 

Jeff Johnson:

These purported reviews, all but one of which is 5 stars, are all in English. Yet none of them is in the US App Store. I also looked at other English-speaking countries such as Australia and Canada but found no reviews. Also suspicious is that the name of every reviewer was a first name followed by an initial for the last name. If you’ve ever looked through App Store user reviews, it’s unusual to find this naming format. In most cases, the name on App Store user reviews is a uniquely chosen username, with no space character, often suffixed by a number.

The final and most conclusive evidence that these reviews are fake, totally invented out of thin air by the developer (or more likely, by an LLM), is that the dates on the reviews, such as “Mar 12,” predate the release of TabControl Extension in the App Store!

God only knows what this obvious AI slop scam is doing with the data from users’ Safari tabs. It’s currently the 12th or 13th most popular Safari extension in the App Store.

Reed Jobs Tells a Neuroscience Story 

Watch this. No words needed, everyone can feel it. Genes are a hell of a thing.

Nature Is Healing 

Basic Apple Guy:

Apple just dropped the hardest icon glow-up for the new Chess icon in macOS 27 Beta 5.

Comparison of the Chess app icons in MacOS 27 Golden Gate: betas 1/2, betas 3/4, and an all new-icon for beta 5. The earlier ones show a flat brown knight. The new one in beta 5 shows a fully-articulated 3D silvery knight. The new one looks great. The previous ones look like shit.

This is not just a tweak or small improvement. This is the difference between total shit and a really nice icon.

The Economist: ‘How to Spot AI Writing’ 

The Economist (‘twas a gift link, but alas, I guess gift views have been used up — here’s an archive link in case the gift link is vexing you):

You can discover AI’s hallmarks by comparing the writing of man and machine. To do this you need a baseline that is distinctive and familiar. The Economist turned to prose that we’re sure is human and that readers will recognise: our own. We designed a study to ask top LLMs — OpenAI’s ChatGPT, Anthropic’s Claude, Google’s Gemini and xAI’s Grok — to write versions of our articles without consulting the web. (As a prompt, we gave them the AI-generated summaries that we have experimentally added to some of our articles.)

This gave us a corpus of human and AI creations and we compared them across 55,940 sentences and 1.2m words. To make sure we were detecting AI quirks rather than our own, we also checked the AI texts against journalism from CNN, the New York Times and the Washington Post. Excerpts from hit novels published between 1950 and 2022 offered another test.

Our findings are surprising. AI prose is distinguishable by word and punctuation choice as well as sentence and paragraph structure. But its hallmarks are not what you might expect, partly because its writing style has changed with software updates. That does not mean that LLMs are great writers: their prose lacks lucidity and elegance and is often formulaic. So those aspiring to be impressive (human) storytellers should avoid the following peculiarities in their own prose.

On point for today.

Some specific findings:

Much of this language could be described as what George Orwell called “pretentious diction”. He railed against writers who “dress up simple statements” with complicated words and jargon to sound clever. Such pontificating penmen, Orwell observed, also believe that “Latin or Greek words are grander than Saxon ones”. (Bots agree: more Latinate suffixes crop up in their writing than in human texts.)

Then look at punctuation. Many believe LLMs stuff their prose with em-dashes, but that is not true after the most recent updates. Today only Claude uses more em-dashes than human writers, with ChatGPT using markedly fewer than any other writer in our study. Humans rejoice — and start using dashes again.

A better way to spot AI-generated writing would be to look for texts without much punctuation at all. LLMs are very Joycean about it: they use fewer commas and semicolons than humans (and hardly any parentheses). They use less punctuation in part because they write longer sentences — “and” is their most overused word — and in part because they do not quote experts.

Alex Micek on BMW’s iDrive ‘Special Surprises’ 

Alex Micek:

In the future, I will feel angry if my car strands me on the side of the road.
I will feel angry if the motor brushes fail expectedly.
I will feel angry if it leaks water.

By contrast, I don’t feel anger when the car tricks me into seeing an advertisement. Instead, I feel a deep, smothering, let-down at the promise of futurism. Bottomless despair is probably inappropriate; a symptom of the misanthropy I should grow out of. But. In the meantime, I feel helpless, discouraged, deeply sad.

Such a good take. I should not let 19 years pass again between linking to his writing.

BMW Probably Paid Sony for the Spider-Man Dashboard Ads, Not the Other Way Around 

This didn’t occur to me but should have — BMW paid Sony to place BMWs throughout the new Spider-Man: Brand New Day movie. They’re not just in the movie but part of the story. So it’s probably the case not that Sony paid BMW to put Spider-Man animations in the dashboards of every late model BMW on the road, but that BMW offered to do it as part of the promotional deal they paid Sony. So BMW probably paid Sony for something that ultimately infuriated a bunch of BMW’s own customers. Geniuses.

Mark Zuckerberg Posts 6,500-Word AI Essay 

Two weeks ago Zuckerberg published a 1,000-word essay, “The AI Future Is for Everyone”. This week he published an expanded 6,500-word version of the essay on Meta’s website. Somehow it almost says nothing more than the original. This tidbit, however, struck me as notable:

For example, in Richland Parish, Louisiana, where Meta is building a large data center, teachers received a $50,000 bonus this year because of the increased tax revenue from our investment. The superintendent told us that teachers are now moving there from across the country and he believes it will become one of the nation’s best school districts.

Unlike the concern about knowledge work displacement, there is a shortage of skilled tradespeople like carpenters, electricians, and construction workers to support the demand for infrastructure buildout. Meta has created America’s Workforce Academy to provide free training for skilled tradespeople and guaranteed high-paying jobs in the areas we’re building data centers.

Those $50K bonuses to public-school teachers are real, and they’re more than most of those teachers’ annual salaries. That’s the sort of thing that could turn public sentiment on data center construction around. But it hasn’t turned around yet.

The Talk Show: ‘Getting the Snack Right’ 

Chance Miller returns to the show. Topics include iOS 27’s progress over summer, live baseball in Vision Pro, the RAM crisis, Apple’s trade secret lawsuit against OpenAI and io, and the EU/DMA situation.

Sponsored by:

  • Squarespace: Save 10% off your first purchase of a website or domain using code TALKSHOW.
  • Factor: Healthy eating, made easy. Get 50% off and one free breakfast item per box for one year, with code talkshow50off.
  • Notion: The collaborative AI workspace where teams and agents work side by side, with a Developer Platform teams can build on.
Netflix Has Peaked 

Andrew Sharp, writing at Sharp Text last month after Netflix’s earnings, asking “Is Netflix Washed Now?”:

I offer this observation as a swirl of heightened anxiety surrounds the company, so let me clarify one thing up front: I’m not predicting imminent doom. Netflix content reaches a staggering 85% of American viewers and has 325 million subscribers globally. Growth is slowing, but that’s the law of large numbers. If practically everyone in America and much of the world is already subscribed to some version of Netflix, and churn rates are still low, then any concern is relative. Going forward: cable is still dying, and even if the biggest premium distribution platform in the world can’t make great content of its own, it can still license movies, TV and sports rights. Netflix can then spread those costs across hundreds of millions of subscribers and a steadily growing ads business, seeing more engagement in a week than Apple TV sees in a year.

So no, the company’s not doomed today or destined for collapse tomorrow. Instead, I think what’s interesting to consider is that Netflix has almost certainly peaked. As a cultural force, as a business success story, and as an entertainment death star destined to swallow Hollywood whole, the arrows are all pointing the wrong direction.

Sometimes the problem with a company isn’t that it isn’t a good business, but merely that it isn’t as good a business as it holds itself up to be. A billion-dollar business ought to be a good thing, but it’s a disaster if the company is valued by investors as a trillion-dollar business. That’s Netflix.

To me it’s backwards that Netflix is valued at $300 billion, roughly the 50th most valued company in the world, and Disney is valued at $180 billion, roughly 120th. Netflix has a streaming video service and owns some IP, like Stranger Things and Squid Game. Disney owns two streaming video services (Disney+ and Hulu), and owns some IP. Little stuff like Mickey Mouse, Star Wars, Marvel, the Disney princesses, and everything from Pixar. Netflix has started dabbling in sports; Disney owns ESPN. Disney owns the world’s biggest and best theme parks and an entire cruise ship line.

It’s true that Netflix’s streaming service is by far the biggest in the world, and likely will remain so for the foreseeable future. But Disney+ combined with Hulu is clearly in second place globally.

For the most recent 12 months:

DisneyNetflix
Revenue$97 B$48 B
Net Profit$11 B$14 B
Net Margin11.5%28%

Netflix is generating slightly more profit on half the revenue. With its theme parks, resorts, and cruise ships, big parts of Disney are effectively hardware, not software — so it’s almost inevitable its margins will be lower than a pure software play like Netflix. But those “hardware” businesses diversify Disney. How confident are you that Netflix will remain the undisputed leader in streaming 10 years from now, or 25? I’d bet good money that Walt Disney World will be the world’s biggest, best, and most profitable theme park 50 years from now. And it’s a joke to compare the value and longevity of the two companies’ IP.

I don’t know that Disney is undervalued by the market, but I sure think Netflix is overvalued. And as for quality, I’ve watched a bit more Netflix original content this summer than I had of late. David Attenborough’s A Gorilla Story was a splendid documentary, but nature documentaries don’t pay the bills. A few other things I watched (or at least watched part of) were absolute dreck. Not high-quality fun trash like Tiger King but amateur-hour I’m-worried-I-lost-a-few-IQ-points-by-watching-it trash. We pay $27/month for Netflix premium and I can say unequivocally that it’s not worth it. I continue paying for it because I’m willing to overpay for it the same way I am for beverages at a hotel — it’s money I’m willing to squander, knowing full well I’m squandering it.

It feels to me like Mr. Market is getting increasingly skeptical whether Netflix justifies its valuation, and Netflix is getting a little squirrelly and desperate trying to come up with an answer. Simple question: what’s their moat? It sure isn’t content quality.

Anthropic Posts ‘How Claude Marks AI-Generated Content’ Without Explaining How Claude Marks AI-Generated Content 

Anthropic support page:

Anthropic has signed the EU AI Act’s Article 50(2) Code of Practice on Transparency of AI-Generated Content, as a provider of both generative AI models and generative AI systems. This article describes how we’re planning to put those commitments into practice, how marking works, and what its limitations are. We’ll update this article and publish more detailed technical guidance as it becomes available.

Regarding generated plain text output:

When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.

Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from. [...]

We’re also working to enable users and other third parties to detect Claude’s embedded watermarks and provenance metadata. Detection checks whether a piece of text or a file carries a supported Claude mark. If a supported mark is found, it indicates that the content may have been processed by Claude.

We’ll share details on detection mechanisms in forthcoming technical documentation.

Anthropic is claiming this is for compliance with an EU law but that “Marking will apply to output from supported models wherever Claude is offered, worldwide.”

This is infuriatingly opaque. I won’t speak to the “signed provenance metadata” they say they’ll be embedding in generated files like PDFs, SVGs, PNGs, and JPEGs. That’s important but it’s complex. Plain text is simple. A string of text is just one character followed by another. At first I presumed that Claude is going to start “embedding” invisible characters/byte sequences between visible characters — which would unavoidably cause immediate problems.

Let’s say I have Claude generate the phrase “Hello world”. (Two words is surely too short to bother with “AI generated” detection — I’m using a two-word phrase here for simplicity in the example.) Claude creates the string of characters “Hello«invisible characters that comprise a “watermark” here» world”. I paste this on the web. Humans who look at it will see “Hello world” but the invisible characters are there. Character-count-limited social media platforms will show that the string contains more than the 11 visible characters in “Hello world”.

If the “watermark” is not comprised of invisible characters but rather visible ones, how in the world does this jibe with their claim that it’s “imperceptible”? This Reddit thread claims that’s how it will work — “It’s a form of steganography where the model subtly biases its word choice to create a statistical pattern that can be detected later.” How in the world can that be squared with “it doesn’t change the meaning, quality, or readability”? And any sort of semantic detection like this is going to cause false-positive problems.

If I ask a tool to generate text or suggest text for me, I expect that tool to generate the best possible word choices it can. Not corrupt its output for the sake of watermarking. If they go through with this, this ought to be poison to the Claude brand.

What happens if you copy my text and quote it in something you write, by hand? Now you’ve got a watermark in your prose, invisible characters or something, that implicates that your prose is AI-generated? Even though you didn’t use AI and just copy and pasted text from me? Madness.

They obviously need to explain exactly what they’re embedding in text if anyone is actually going to detect it, but if they explain it, anyone can simply remove it. And what happens, for example, if someone just OCRs Claude-generated text rather than selects and copies it? This is all so stupid. I have never been happier that I’ve never actually used Claude for anything.

What’s New in iOS 27 Beta 5 

Zac Hall at 9to5Mac has a good rundown of new features and changes in beta 5. Some app icons got tweaks (the Siri app in particular looks much cooler now), and a bunch of the existing American and British Siri speaking voices now have the Pace and Expressivity adjustment sliders.

Also interesting: the OS 27 releases were stuck at beta 4 for 22 days. That’s a week longer than usual. Will Hains (my Kotoba friend) maintains a wonderful single-page website tracking iOS version release date history. For these summer “next big version” comparisons, it’s best to filter by “*.0” releases. Do that at Hains’s website and you can see a fairly clear pattern: betas 1 through 4 typically last about two weeks each; then starting with beta 5, Apple switches to a roughly weekly schedule until the end of the beta cycle. I wouldn’t read anything into that other than to expect beta 6 next week and beta 7 the week after that.

‘Friday Night Baseball’ Will Start Broadcasting Games Live, in Apple Immersive, on Vision Pro 

Apple Newsroom:

For the first time, baseball fans can take in all the action of “Friday Night Baseball” live in Apple Immersive on Apple Vision Pro, leveraging 3D video recorded in 8K with a 180-degree field of view. On August 28, viewers with Vision Pro will experience the iconic Red Sox vs. Yankees rivalry from unique camera angles that place viewers right in the game, immersed in the drama of each at-bat, the celebrations in the dugout after a big hit, close plays at the outfield wall, and more.

The experience features bespoke commentary from Paul Severino (play-by-play), Xavier Scruggs (analyst), and Michelle Margaux (sideline reporter); dedicated replays; and an immersive graphics package that includes team lineups and stats floating in the viewer’s space. During breaks in the action, the broadcast remains live, immersing viewers in the stadium experience from all angles as ambisonic microphones capture sounds in Spatial Audio.

Hot damn, they even picked the best rivalry in sports to kick it off.

Michael Tsai on My Retraction of the Astrology/Astronomy App Store Rejection Story 

Michael Tsai:

I’m really unhappy about this, both for my role in spreading a false story and because I think it will hurt the cause of reforming App Review. I know there are many crazy rejections and have experienced some first-hand. This story was believable because of that well-known history and because Godier didn’t seem like a nobody trying to get attention — he was the developer of another highly regarded app. But now people are going to point to true developer stories and accuse them of being fakes, too.

Same.

‘The Problem With Vibe-Coded Flattery’ 

Ernie Smith at Tedium:

I’m seeing email after email suggesting these users built a thing they’re truly passionate about. But if we can’t trust that this passion is real, legit, and from the heart, we’re in trouble.

Take a step back. Is this thing you built a reflection of you?

I think it’s fun and exciting that I’m getting a lot more emails about a lot more new apps. But a fair chunk of those emails are themselves clearly written by an AI chatbot. It’s so obvious. Those go straight in the trash.

The NYT and WSJ on Apple, China, and the RAM Crisis 

Kalley Huang, reporting from San Francisco, and Ana Swanson, from Washington, for The New York Times (gift link):

Apple, which has cited the memory chip shortage for recent price increases, is working with other consumer electronics companies to push for permission to buy memory chips from China, which faces certain restrictions, seven people said.

Apple was in talks in 2022 to buy memory chips from Yangtze Memory Technologies Corporation, or YMTC, but those efforts fizzled under pressure from U.S. lawmakers.

Many officials appear unsympathetic to Apple’s desire to buy Chinese chips, because the company has promised for years to move more of its supply chain to the United States but has made only small steps toward that goal, four people said. Other critics of the plan said that China was also experiencing a similar chip crunch, and had proved itself to be a risky supplier when it restricted its mineral exports last year.

Raffaele Huang and Rolfe Winkler, reporting for The Wall Street Journal (also a gift link):

Apple has been testing memory chips from China’s CXMT across product lines including iPhones and MacBooks, according to people familiar with the matter, as the U.S. company addresses a memory crunch during the artificial-intelligence boom.

Apple has held early talks with CXMT about supplying components with the goal of using them in some devices sold in China, the people said. Apple hopes to win the White House’s blessing to do business with the Chinese company. [...]

Apple often tailors components to its devices to maximize performance. Using standard CXMT memory chips could force Apple to redesign certain parts of its products. Industry analysts also say CXMT’s technology still trails its foreign competitors, in part because the company’s access to the most sophisticated chip-making equipment is restricted.

At present, CXMT wouldn’t be a panacea for Apple. The Chinese company has maxed out its production capacity this year, leaving little room for new international clients, the people familiar with the matter said.

The gist of both reports is that just like in 2022, Apple might not get permission from the U.S. government to use Chinese RAM, even for devices sold in China. And, even if they do secure permission, it might not actually help much with the supply crunch.

WorkOS: Connect Your Agents to Your API 

My thanks to WorkOS for, once again, sponsoring DF last week. What’s the best way to connect AI agents to your API? REST is great for human developers; MCP serves the agents. Many teams treat REST and MCP as competing standards, and through choice or necessity, pick one. But they ought not be considered rivals. They’re separate layers. Most MCP servers just call REST internally to do the real work. The best ones don’t convert every endpoint into a tool, they focus on what the agents are actually trying to accomplish.

That layered approach also means shipping OAuth 2.1 with scoped tokens. WorkOS AuthKit already speaks that spec, so you skip building an auth provider on top. Check out WorkOS’s breakdown on MCP vs. REST.


Retraction: The App Store Rejection of the Week That Was, in Fact, a Correct Rejection

Yesterday I published an article titled “App Store Rejection of the Week: Dark Hours”. I have retracted it. Its premise was so fundamentally wrong that there’s no point merely correcting or editing it. Even the title, as I explain below, was inaccurate. Although the original is now retracted, I’m not memory-holing it. The text of the original story is available, for transparency and accountability, in plain text (Markdown, natch) and PDF (preserving original presentation). Both of those versions include a preface at the top linking to this retraction. The URL for the original story now redirects to this one that you are currently reading.

To the best of my recollection, this is the first post I’ve retracted in the 24 years I’ve been writing Daring Fireball. I hope it’s the last. I was misled, both overtly and through omissions, in several ways, but what I publish is my responsibility, and I apologize for the error.

Terry Godier first came to my attention in February, when I linked to his excellent interactive essay on RSS feed reader design, “Phantom Obligation”, which essay introduced Current, Godier’s new RSS reader that he made to exemplify the ideas from his essay. Current is, deservedly, a bit of a breakout hit. (E.g., David Pierce, at The Verge, put it on a very short list of iOS-exclusive indie apps that keep him from switching to Android.) In March I linked to another interactive essay from Godier, “The Last Quiet Thing”, and again in April to a post regarding the App Store’s 5-star review system. I struck up an iMessage correspondence with Godier around when I first linked to his work.

Yesterday Godier posted “Browsers Have Standards, the App Store Has Judgment”. As originally published, that post contained these two paragraphs:

A while ago I tried to submit an iOS app for Dark Hours, my astronomy website for normal people. It was rejected on the grounds that it was astrology.

It has no tarot function, no horoscopes, and nothing that I, or anyone else I’ve asked, would associate with astrology.

Those paragraphs, at this writing (August 8, 11:00 pm ET), have been deleted and replaced by this:

Note: the original version of this post had a section here about an astronomy app I am working on that began as an astrology app and was rejected after having the astrology content removed. App store [sic] review reached out to me and let me know that they had apparently never been given the updated build and that the app should be fine to submit now.

You can now see the problem that has led me to fully retract my original post, given that my post was entirely predicated on the premise that Godier’s original description of the rejected app was true — that the app “has no tarot function, no horoscopes, and nothing that I, or anyone else I’ve asked, would associate with astrology.” The truth is, the app, as originally submitted by Godier to the App Store (under the name “Asterly”, not “Dark Hours”), was entirely dedicated to astrology, not astronomy, and did in fact include a “Tarot card of the day” feature amongst other occultist horseshit.

The grounds of Apple’s original App Store rejection of the app, and the rejection’s upholding by the App Review Board, were correct.1 I wrongly took Godier at his word, both in his public blog post and in private iMessage correspondence yesterday, that the rejection wasn’t just merely debatable, but completely and rather preposterously ungrounded. Whether Godier ever submitted a build of “Asterly” to the App Store that contained no occult horseshit and only the hard-science astronomy features that were present in his “Dark Hours” website that was available for the last week, I don’t know. But I have no reason to believe that he did.

My disdain for astrology is so utter, and my esteem for Godier’s previous work so high, that it simply never occurred to me that he might have actually made and submitted to the App Store an astrology app, let alone that he’d then feign surprise and frustration that an astrology app was rejected for being an astrology app. I showed him a draft of my post before publication, to make sure I had the story straight, and he offered not a word of caution, only gratitude for my drawing attention to the matter.

It gets messier. Godier’s astrology app that he submitted to the App Store back in January was named “Asterly”. That was still the name when the App Review Board upheld its rejection in April. According to Godier, frustrated by the App Store’s rejection, he ported the astronomy version to the web, launching it last week under the name “Dark Hours” at the domain darkhours.io. (This is why it was incorrect for me, in the very title of my post, to claim that Apple had rejected an app named “Dark Hours”. They rejected an app named “Asterly” and had never seen an app from Godier named “Dark Hours”.) Yesterday, after I linked to Godier’s post and Hacker News then linked to my post, Godier’s “Dark Hours” (with a space) came to the attention (and justifiable surprise) of Miguel Beher, creator of an open-source “astrophotography and dark-sky planner” project named DarkHours (no space). Beher’s GitHub project contains the source code, and the actual web app is freely available at the domain darkhours.app. In an uncomfortable exchange between Beher and Godier on Bluesky, Beher pointed out that Godier’s Dark Hours had the same bug as Beher’s that routed people to “random fields in Mexico”. Earlier today, Godier took his web app down and redirected his darkhours.io domain to Beher’s darkhours.app.

At the end of my now-retracted post yesterday, assuming I had righteously and rightly skewered Apple for an egregiously erroneous App Store review rejection, I wrote:

Mistakes happen. But in a functioning system mistakes get corrected, and mistakes as obvious as this one get corrected almost instantly and include a quick apology for the conflation.

Obviously it was I who was mistaken. This article is my correction, and I apologize to all who read and believed my now-retracted post, and to the reviewers at the App Store whose competence (if not literacy) I besmirched. I am deeply sorry about that. Won’t happen again. 


  1. Regardless of one’s opinion regarding pseudoscience and occultist horseshit — pro, con, or indifferent — one might reasonably think it wrong for Apple to disallow or discourage such apps from the App Store. In fact, there exist plenty of such apps in the App Store, and Apple has even run “Best Astrology Apps” editorial features. Apple’s stance is basically that the App Store has enough of these apps (and I suspect they’re a common source of scams, given that by their very nature they target the gullible). Guideline 4.3(b) states (emphasis added):

    Certain kinds of apps, such as dating, flashlight, sound effects, wallpaper, simple timers, and fortune telling, are well established on the App Store and we will not accept new submissions unless they offer a meaningfully different or improved experience. We may remove these apps from the App Store going forward if they are not updated, improved, or do not attract customers. Other kinds of apps, such as drinking games, Kama Sutra, fart, and burp apps, are mediocre, low-quality, or low-effort and do not add value to the App Store.

    That serial comma, ever useful, gives hope to anyone hard at work on a “fart and burp” app. ↩︎


Corrupt Minds Think Alike 

Tariq Panja, reporting for The New York Times:

The five men should have been in a celebratory mood. FIFA had just pulled off a World Cup that broke records on a number of fronts. Its spectacular culmination, the final, was a day away from kickoff.

But their mood was anything but jubilant. The men, a group of top directors known as the FIFA bureau of the management board, had been summoned to a hastily organized meeting at the Waldorf Astoria hotel in Manhattan, according to three people familiar with the discussions speaking on condition of anonymity because the talks were private. The directors had until midnight to sign off on a plan that threatened to change soccer forever and would almost certainly lead to a major schism among the sport’s leaders.

The project purported to sell a 20 percent stake in FIFA to a group of investors led by Joshua Kushner, the brother of President Trump’s son-in-law Jared Kushner. Such a move would put into private hands a portion of the World Cup, an event that has been run and controlled by FIFA, a nonprofit in Zurich, for 96 years.

I don’t know what’s dumber — this stupid corrupt scheme to privatize the World Cup, or me, for not assuming from the start that the Trump family was involved in it somehow.

Maybe ‘Steal Underpants by Blowing a Fortune on AI Tokens’ Is, in Fact, Not a Good Business Plan 

Joseph Cox, writing for 404 Media (paywalled, sans gift links, alas):

Consulting giant Accenture is trying to figure out how to stop non-technical workers from blowing through companies’ AI token budget on trivial tasks like converting PDFs to presentation slides, according to leaked audio obtained by 404 Media. Across the industry Accenture is seeing “soaring token spend,” according to the audio. [...] It also undercuts the narrative that superpowered engineers generating mountains of code are behind the AI boom. In many cases it is non-technical staff burning through tokens for non-specialized tasks. [...]

At one point in the meeting, Kwak and Eduardo Salamanca de Diego, senior manager of product management at the company’s Center for Advanced AI, start presenting about what is described as “token ops.”

Kwak says he knows people aren’t using slides these days, but he has some. As he appears to be preparing to present, Stuart Henderson, Accenture’s client group lead, interrupts. He jokes he hopes Kwak didn’t just convert a PDF into images and then into Markdown files. “I’m learning that’s one of the big token chewers,” Henderson says. “Turning PDFs into Markdown: is that right?”

That’s when Kwak says that’s what Accenture’s own data shows.

My impression of consulting giants like Accenture (McKinsey, Bain, Deloitte, KPMG...) has always been that they are very good at looking smart but in fact very bad at actually being smart. The idea that they’re spending serious money on AI tokens to turn PDFs and PowerPoint decks into Markdown only makes me think I’ve overestimated the median intelligence of the people who work there. I really thought Markdown couldn’t get more successful but this takes the cake. I love it that my little baby is helping burn these companies’ cash. This is so stupid it hurts to think about.

Markdown is something you start with because it’s easy to write and universally readable. You write Markdown and use that to make slides — like, say, with iA Presenter — not the other way around.

(Via Simon Willison.)

Simon Willison on Blogging 

Simon Willison, with Cynthia Dunlop for her tech blogger interview series (from back in January, but he just got around to linking to it so I just got around to seeing it):

Any lessons learned that you want to share with the community?

My number one tip for blogging is to lower your standards! Aim to hit publish while you are still actively unhappy with what you have written, because the only alternative is a huge folder full of drafts and never publishing anything at all.

Nobody will ever know how perfect the thing you intended to write would have been. The flaws you see in your writing are invisible to everyone else.

This is excellent advice. Me, I try to get into the mindset of playing live music, not recording a studio album. Except when I’m writing a piece where I really want it to be an album. Those aren’t rare, per se, but they’re occasional. If I tried to make every post a hall-of-famer I’d never get anything out.

I’m aiming for professionalism. I’m performing live in front of an audience — not just jamming in my garage or bedroom, fucking around. So I’m careful and concentrate. I want to hit every note, in time. But at my best I’m moving from song to song.

But that leads to Dunlop’s next question for Willison:

Your advice for people just getting started with blogging?

Just start. It’s so easy to get caught in the trap of obsessing over the design of your blog, and planning for content that you never actually get around to writing. As long as each entry has a date on it and a permanent URL, it counts as a blog. I think adding an Atom or RSS feed is important too, but you can get started without one — don’t treat that as a blocker.

Don’t worry if nobody reads it. People are unlikely to stumble on your blog organically, but what matters is not the quantity but the quality of your readers. If the only person who reads your blog is a hiring manager that you send a link to, and that gets you an interview, your blog has already paid itself off many times over.

That’s the hard part about getting started, or at least it was for me. When you already have an audience it’s easier to take it seriously, to concentrate, to focus, to strive for perfection even though you will fall short. If you’re playing music in an empty room it’s harder to do that, to play the same way you would with an audience of dozens, hundreds, or thousands of people in front of you.

But just imagine they’re there. They will be if you keep it up. That’s one of my bits of advice to new bloggers: I firmly believe all blogs eventually get the audience they deserve, if they keep going. Write for the audience you want, not the audience you have.

Google Earth Retracts AI Tool for Making Fake Satellite Images After It Was Immediately Abused Upon Release 

Jeremy Hsu, Ars Technica:

Google briefly allowed anyone to create AI-modified versions of satellite imagery available in Google Earth — before quickly reversing its decision as people shared examples of AI-generated pictures that illustrated the potential for misinformation and disinformation.

Momentum!

Some New Data Centers Are Necessary 

A bunch of mine are stored here.

An AI Model From Meta Also Hacked Another Company During Testing 

Simon Willison:

So that’s Anthropic, OpenAI, and Meta. Google Gemini really needs to catch up on accidentally cyberattacking other companies.

I’m surprised Willison hasn’t heard about how much momentum Gemini has.

Meta: Introducing Muse Code and Muse Spark 1.2 

Meta AI:

We’re excited to release Muse Code (beta), a terminal coding agent powered by Muse Spark 1.2, our newest model. This marks our next step toward the frontier, with larger and much more capable models on the way. [...]

Muse Code takes on complex software engineering tasks across large repositories: planning changes, writing code, and validating the results. It can coordinate multiple persistent subagents for each task, solving difficult problems faster, more accurately, and with less intervention.

Terminal-only (so far), so it avoids the whole native-app-versus-Electron-shit-sandwich debate. They’ve embedded some fun interactive examples of Spark’s output right in the web page.

Simon Willison:

An interesting twist on pricing is that the model is offered as two different model IDs. muse-spark-1.2 is priced at $1.25/million input and $4.25/million output — close to Gemini 3.6 Flash ($1.50/$7.50) — but if you agree to let Meta use your data “to improve our products” you can use muse-spark-1.2-contributor which is $0.10/$0.20 — a huge discount, closer to GPT-5.6 Luna ($0.20/$1.20) and Gemini 3.1 Flash-Lite ($0.25/$1.50).

That’s over a 10× discount in exchange for giving your data to Meta. I can see the appeal if you don’t care about the privacy of your data and code, like, say, if you’re using Muse Spark to generate code for a throwaway project, or for something that’s open source anyway. But I wonder if Meta has considered that this offer might spook would-be users who absolutely do not want to grant Meta the rights to their data. One would hope that the two systems have a strong firewall between them. But given Meta’s well-established contempt for the sanctity of user data, it’s not at all unreasonable to suspect that the difference here is like reserving a no-smoking seat in an airplane with a smoking section.


OpenAI Responds to Apple’s Lawsuit and Motion for Preliminary Injunction: ‘Apple Is Getting This Wrong’

OpenAI published an unbylined blog post overnight, responding in public — but not yet in court — to Apple’s new motion for a preliminary injunction. It’s an unusual move to respond to a high-stakes legal filing with a blog post, but OpenAI is an unusual company. A few snippets from their post, and some commentary:

Apple had claimed that they contacted OpenAI in February and that we didn’t respond. They now admit that their outside lawyers emailed the wrong person after confusing two Asian last names — only after we brought this to their attention.

OpenAI is hanging on to the fact that Apple’s outside counsel, Gabriel Gross, sent one email to the wrong address, and quickly emailed an apology. In OpenAI’s phrasing, it sounds like Apple’s attorney sent the entire initial letter of concern to the wrong person, and that’s why OpenAI never responded — because it wasn’t sent to the correct person (OpenAI general counsel Che Chang). That’s not what happened. The initial blockbuster “hey we think you guys are stealing our trade secrets and we want to talk to you about it” letter was sent to Che Chang. And Che Chang never did respond to Apple’s lawyers. That a mistaken email thanking Che Chang for a phone call that never happened (because that email was intended for another OpenAI employee) was also sent is irrelevant. I don’t understand why OpenAI is continuing to focus on this inconsequential mistake. (Apple’s motion for a preliminary injunction includes the full text of the mistaken email and subsequent apology.)

Apple accuses Chang Liu of accessing Apple confidential information after leaving the company, but only now admits that Apple employees reached out to him and asked for his help to locate this information (you can read the messages here). Apple now tries to shift the blame to “residual access”, but they also don’t disclose that this is a common issue with Apple which is caused by them failing to properly manage system access when people leave. What that means in practice is that former employees who are trying to do the right thing when they leave still have access to Apple files — despite not wanting them or even being aware of them.

OpenAI is seemingly alluding to Apple’s unusual use of iCloud Drive, tied to employees’ personal Apple Account IDs, that I (coincidentally?) wrote about yesterday. Apple’s motion for injunction, however, addresses this very point. From page 3 of the motion:

Mr. Liu resigned on Thursday, January 22, 2026, and provided notice that he would start at OpenAI the following Tuesday. On his last day, he failed to respond to Apple’s attempt to schedule his exit interview or sign his confidentiality reminder.

In the days following his departure, Mr. Liu seemed initially cooperative and aware of his obligations to Apple. He worked with others on his former Apple team to return certain Apple information remaining on his personal iCloud account to Apple.1 He also continued to converse with former co-workers, for example, to answer questions about his earlier work and where certain information was stored. But these interactions and exchanges cannot explain the repeated, unauthorized downloading of voluminous technical files from Apple’s cloud-based storage discussed below, which Mr. Liu performed on multiple occasions from February to April 2026 while employed by OpenAI.

That footnote reads:

1 While Apple seeks discovery into what Apple confidential information Mr. Liu accessed from his personal storage accounts (including iCloud) and devices after his departure, the specific unauthorized downloads referenced in the complaint and at the heart of this motion are not based on iCloud activity, but instead relate to Apple’s third-party cloud storage.

Nowhere in any of Apple’s filings (here’s the Court Listener index page for all the documents filed in the case) does it say who the third-party cloud storage provider is, but I’m almost certain it’s Box, which I know is widely used throughout Apple.

The iMessage transcripts that OpenAI provides at the bottom of their post do not contradict Apple’s claims at all. Apple’s motion states that Liu helped former colleagues find certain documents that were in iCloud; that’s what OpenAI’s transcript shows. But that’s not in dispute. Apple also claims that Liu accessed confidential information, presumably in Box and definitely not in iCloud Drive, on five different occasions, up until 27 April 2026, over three months after he left Apple. These chat transcripts offer no explanation for that. The chat transcripts explain iCloud Drive access that Apple itself says is not in dispute, and do not explain the 37 documents Liu downloaded from the third-party cloud provider (Box?) that Apple says are at the heart of naming him in the lawsuit. Here is Apple’s declaration from digital forensic specialist Daniel Roffman, documenting Liu’s access to confidential files post-employment (albeit with significant redactions).

I do not understand why OpenAI is treating this as a PR problem instead of as a legal problem. Dan Moren, linking to it from Six Colors, is of similar mind, writing:

What kept running through my head while reading this was the old legal chestnut: “If you have the facts on your side, pound the facts. If you have the law on your side, pound the law. If you have neither on your side, pound the table.”

Thus far this feels like table-pounding from OpenAI to me. Their blog post does, however, move the ball from “we have no interest” in Apple’s trade secrets to “we don’t have them”, (emphasis added):

Apple also accuses Tang Tan of trying to get and use their trade secrets. However, Tang has always been clear with the team that we do not want, and must not use, any confidential information from other companies. Tang served Apple for more than 24 years and was widely known as one of the most innovative leaders at the company. [...]

Apple’s request for a preliminary injunction is both based on false information and completely unnecessary because we do not have, nor want, any of their trade secrets. We’re much more interested in building innovative products and technologies that push the frontier.

To me, the most interesting response from OpenAI wasn’t their blog post. It was an email released by Apple, as “Exhibit F” to one of their expert declarations submitted to the court last night. OpenAI has retained the renowned law firm Quinn Emanuel as outside counsel, and this exhibit is a long email from Quinn Emanuel attorney Patrick Curran to Apple’s attorneys. From that email, dated Monday July 20, Curran writes:

You also ask that we “revisit” the specific points proposed in your July 15 letter. It appears that you want to move backwards. As noted, we already discussed these during our meet and confer but Apple was unable to respond to basic questions my colleagues raised about these requests. For example, your letter proposes that OpenAI “[p]roduce witnesses to testify at deposition” but Apple was unable to identify who those witnesses would be. Similarly, Apple was unsure when we asked if it was actually proposing that hundreds of OpenAI employees fill out “questionnaires” even if Apple has no basis to allege (and is indeed not alleging) that such employees have any connection to this litigation. The seven sections in your letter are broadly worded and remain vague and general. This is not what a forensic protocol looks like and we’re sure you understand that you will not get this as relief from the court. You first need to (preliminarily) identify the TS you are suing for, and your email states that you “appreciate the need” to do so. Any protocol will be informed by such identification. A forensic protocol cannot be based on general terms like “Apple confidential information”; you need to tell us what you’re looking for, and it sounds like you understand that and are prepared to do so. The efficient way forward is therefore to tackle these issues as part of the negotiation of a proper, detailed forensic protocol. If you instead prefer to move for a PI because OpenAI did not agree off the bat to subject hundreds of employees to “questionnaires” about “Apple confidential information” generally, that is unfortunate — and inconsistent with what I understand both our clients have requested. If you choose this path instead of working with us, we look forward to filing an opposition that sets the record straight.

Apple, obviously, did choose this path (“PI” = preliminary injunction), and I too look forward to OpenAI’s setting the record straight, especially if they do so in plainspoken language like Curran’s in this email. Curran continues:

Finally, although I know OpenAI would like to resolve this amicably, as their counsel I have to tell you what I think you already know — this case lacks merit. You have not articulated any basis to support a preliminary injunction. Your complaint is predicated on a misrepresentation of facts and allegations that are speculative at best. It fails to even remotely identify any trade secrets. You are attacking ordinary business practices (used widely across the industry). You are complaining about situations that you have caused, including through your own procedures and decisions. We stand ready to oppose any preliminary injunction motion and tell the world what really happened here to set the record straight. We made clear we would prefer to quickly and collaboratively address any legitimate concerns that your client has, but that is not well-served by repeated threats.

This email is a far better response than what OpenAI published on their blog.